Papers with voice conversion
Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in speech synthesis have improved the quality of polyglot voices. |
| Approach: | They propose a cross-lingual any-to-one voice conversion system that preserves the source accent without multilingual data from the target speaker. |
| Outcome: | The proposed system preserves source accent without multilingual data from target speaker and reduces training data requirements. |
Eta-WavLM: Efficient Speaker Identity Removal in Self-Supervised Speech Representations Using a Simple Linear Equation (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for learning meaningful representations from unannotated data are resource-intensive and degrade other speech components. |
| Approach: | They propose a method that decomposes SSL representations into speaker-specific components and generates speaker disentangled representations. |
| Outcome: | The proposed method achieves speaker independence and improves on state-of-the-art methods. |
Playing with Voices: Tabletop Role-Playing Game Recordings as a Diarization Challenge (2025.findings-naacl)
Copied to clipboard
| Challenge: | Using a small dataset, we propose that audio of tabletop role-playing games (TTRPGs) could serve as a challenge for speaker diarization systems. |
| Approach: | They propose that audio of tabletop role-playing games (TTRPGs) could serve as a challenge for speaker diarization systems. |
| Outcome: | The proposed system can pick the speaker and determine that impersonating is just that. |
SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing (2022.acl-long)
Copied to clipboard
Junyi Ao, Rui Wang, Long Zhou, Chengyi Wang, Shuo Ren, Yu Wu, Shujie Liu, Tom Ko, Qing Li, Yu Zhang, Zhihua Wei, Yao Qian, Jinyu Li, Furu Wei
| Challenge: | Existing work shows that pre-trained models can improve in various natural language processing tasks. |
| Approach: | They propose a unified-modal encoder-decoder framework that pre-trains speech-text representations using large-scale unlabeled speech and text data. |
| Outcome: | The proposed framework is superior to existing models on speech-to-text processing tasks. |
Voice synthesis in Polish and English - analyzing prediction differences in speaker verification systems (2025.coling-main)
Copied to clipboard
| Challenge: | Using audio deepfakes, we can create high quality false voice recordings convincing enough to deceive human ears and pose security concerns. |
| Approach: | They examine the effects of deepfakes on speaker recognition systems across English and Polish corpora, evaluating both Text-to-Speech and Voice Conversion methods. |
| Outcome: | The proposed methods can maintain personal traits, posing risks of unauthorized access, and can be used to deceive human ears. |
O_O-VC: Synthetic Data-Driven One-to-One Alignment for Any-to-Any Voice Conversion (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Traditional voice conversion methods attempt to separate speaker identity and linguistic information into distinct representations, but this method often leads to information loss during training. |
| Approach: | They propose a method that leverages synthetic speech data generated by a pretrained model . synthetic data pairs that share the same linguistic content are used as input-output pairs . |
| Outcome: | The proposed method outperforms state-of-the-art methods in speaker-to-voice conversions. |
A Unified Feature Mixture Framework for Joint Speech and Singing Deepfake Detection (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for deepfake detection fail under speech-to-singing domain shift . a speech-retentive multi-domain fine-tuning strategy enables adaptation to singing . |
| Approach: | They propose a unified deepfake detector based on a multi-branch mixture-of-experts architecture that integrates three complementary feature views. |
| Outcome: | The proposed detector achieves 1.82% EER on CtrSVDD, compared to 37–62% for existing detectors . it can generalize to unseen generators and preserve strong speech performance . |